Genome Biology
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Genome Biology's content profile, based on 637 papers previously published here. The average preprint has a 0.47% match score for this journal, so anything above that is already an above-average fit.
Daito, Y.; Uechi, M.; Kinoshita, T.; Tonosaki, K.
Show abstract
Background: Accurate identification of differentially methylated regions (DMRs) is fundamental to epigenomic research but remains challenging due to biological variability among replicates, heterogeneous effect sizes, and the tendency of adjacent cytosines to share similar methylation states. Many existing methods aggregate methylation measurements before statistical testing or do not explicitly account for replicate-level variability, contributing to elevated false-positive rates. Results: We developed glmmDMR, a DMR detection framework that combines generalized linear mixed models with a seed-based strategy for reconstructing DMRs from locally high-confidence signals while explicitly modeling replicate-level variability. Using simulated datasets with known ground-truth DMRs, we demonstrate that false-positive detections are more strongly associated with methylation variance among biological replicates than with the magnitude of methylation differences between groups. glmmDMR achieved higher precision than existing approaches while maintaining competitive recall, particularly for subtle methylation differences. Site-level modeling with beta regression provided the strongest overall performance, and seed-based region construction reduced artificial DMR fragmentation, improving recovery of true DMR boundaries and producing more contiguous, biologically interpretable DMRs. Applied to Arabidopsis thaliana ddm1 methylomes and a rice DEMETER-LIKE DNA demethylase mutant (Osdml3a-1), glmmDMR identified biologically meaningful DMRs, revealing widespread TE-associated hypomethylation and subtle TE-family-specific hypermethylation. Conclusions: Replicate-level methylation variance is an important determinant of DMR detection performance, and explicitly modeling this variance improves discrimination of biologically meaningful methylation changes from high-variance signals. By combining variance-aware statistical modeling with seed-based region construction, glmmDMR provides a robust framework for identifying contiguous, biologically interpretable DMRs across diverse methylome datasets.
Ardaman, A.; Forgiarini, C.; Arunkumar, R.
Show abstract
Intraspecific hybridization in allopolyploid plant genomes has the potential to induce non-additive changes in gene expression and DNA cytosine methylation, partly through interactions among divergent parental subgenomes. However, the extent to which intraspecific hybridization reshapes gene expression, coordinates homoeolog regulation, and remodels methylation in higher-order polyploids remains poorly quantified. To address this, we sequenced seedling leaf transcriptomes and methylomes from two parental cultivars of hexaploid bread wheat (Triticum aestivum L.) and their hybrids. More than 40% of genes were differentially expressed between hybrids and parents, although many were not differentially expressed between the parents themselves, consistent with complex trans-regulatory effects in the hybrid genome. This effect was more pronounced for homoeologs whose relative expression differed between the parents. These expression shifts often occurred simultaneously across all three homoeologs within triads, reducing homoeolog expression bias (HEB) in the hybrids. CG methylation levels were similar between the parents and hybrids in regions of low genetic divergence and in transposable element (TE)-rich regions, whereas CG sites in gene-rich regions showed more additive inheritance (hybrids intermediate between parents), particularly when parental haplotypes were themselves divergent. TE and gene body methylation (gbM) was strongly conserved in parents and hybrids. gbM was associated with more balanced homoeolog expression and fewer non-additive expression changes. CHH methylation showed overdominance, whereas non-conserved CHG methylation was enriched in TE-rich regions, suggesting that non-CG remodeling may reflect parental differences in TE and small-RNA content. Our results show that intraspecific hybridization within a hexaploid species can generate non-additive changes in gene expression and DNA methylation in seedling leaf tissue, while the presence of homoeologous genes, parental HEB, parental genetic and methylation divergence, and genomic location have varying levels of influence on expression or methylation remodeling.
Runyan, M.; Gupta, S.; Leshaem, Y.; Geller-McGrath, D.; Liu, C.; Saferali, A.; Dy, J.; Radivojac, P.; Tesfaigzi, Y.; Castaldi, P.; Paul, A.
Show abstract
BackgroundAlternative splicing, the mechanism by which intronic sequences are excised from pre-mRNAs to produce mature mRNA, affects >95% of human protein-coding genes and is a major driver of human disease states. The spliceosome, a protein-RNA complex responsible for splicing pre-mRNA, identifies candidate splice sites partly through the recognition of characteristic sequence motifs at exon-intron junctions. Deep learning models that predict the presence of splice sites from pre-mRNA sequence have achieved breakthrough performance relative to previous machine-learning techniques, and these models have improved our ability to identify pathogenic genetic variants that alter splicing. ResultsWe show that, while overall performance measures from these models suggest near-perfect performance, substantial gaps in prediction remain, including the identification of splice sites with low usage rates and tissue-specific splice sites. We leverage one of the largest paired RNA and genotyping datasets used to date to train a novel splicing model optimized for a specific cell type, human airway epithelial cells. We trained a dilated convolutional neural network on data from cultured airway epithelial cells from 100 donors, and showed that this model outperforms current state-of-the-art models on splice site identification and splice site usage quantification, including on multiple tissues not included in the model training data. ConclusionsWe present the most comprehensive evaluation of state-of-the-art splicing models published to date, revealing reasonable performance across models for genetic variant effect prediction along with important performance gaps and insights into directions for future model development.
Ebler, J.; Prodanov, T.; Blair, A.; Lee, S. K.; Ebert, P.; Human Pangenome Reference Consortium, ; Paten, B.; Marschall, T.
Show abstract
Pangenome graphs built from haplotype-resolved de novo assemblies enable accurate analysis of genetic variation. The short-read-based tool PanGenie efficiently genotypes variants discovered in a pangenome across large cohorts and outperforms linear reference-based methods for structural variants (SVs). However, it cannot detect novel variants absent from the graph, missing many rare SVs (allele frequency <1%) and was limited to graphs with 254 haplotypes. First, we introduce a haplotype sampling step that reduces the number of haplotypes using sample-specific k-mers before genotyping, decreasing runtime twelvefold and memory usage 1.4-fold at 30x coverage. Second, we present a polishing workflow that corrects residual errors in haplotypes inferred from PanGenie genotypes and incorporates rare and private mutations. We genotype 3,202 samples from the 1000 Genomes Project and use low-coverage ONT data (967 samples) for polishing. We achieve a median QV of 46 and provide the 1,934 polished haplotype sequences as a community resource.
Yin, R.; Li, D.; Zong, W.; Ketchesin, K. D.; Seney, M. L.; McClung, C. A.; Baldoni, P. L.; Tseng, G. C.
Show abstract
Correcting for library size is an essential step in bulk RNA-seq analyses, as differences in sequencing depth across samples can obscure biological signal with technical noise. While numerous normalization methods and model-based strategies have been proposed, we demonstrate here that library size-normalized counts and differential expression results obtained from such widely adopted approaches often remain strongly correlated with library size in large-scale RNA-seq experiments. Through a systematic analysis of over 100 publicly available GEO and TCGA RNA-seq datasets with raw count data, we show that library size association is observed for a substantial proportion of genes even after state-of-the-art library size correction approaches recommended by leading normalization tools. To address this issue, we propose gecco, a gene-specific exponent-corrected normalization method for RNA-seq counts that incorporates library size directly into the statistical framework via a gene-specific correction term, rather than applying a uniform adjustment factor across all genes. This formulation generalizes existing normalization approaches and yields normalized counts that are free of residual library size effects. Using both simulation studies and real large-scale RNA-seq datasets, we show that our method mitigates library size bias while preserving biological signal across a range of parameter settings. We further demonstrate that our approach leads to higher detection accuracy and more biologically meaningful pathway enrichment results in downstream differential expression and rhythmicity analyses without compromising false discovery rate control. Our method is implemented in R and is fully compatible with the widely used differential expression analysis methods DESeq2 and edgeR.
Yeung, J.; Tan, J.; Wang, L.; Wu, D.; Melo Carlos, S.; Kageyama, J.; Kamm, J.; Chu, B. B.; Mayba, O.; Forrest, W. F.; Xie, S.
Show abstract
Perturb-seq measures transcriptomic responses to genetic perturbations at scale, but conventional designs that enrich for one guide RNA per cell remain resource-intensive. Standard analyses discard cells carrying multiple guides, further limiting the usable yield from each experiment. Here, we characterize how incorporating these guide multiplets affects signal recovery, information loss, and cost reduction. At the highest guide burden, cells showed increased stress and suppressed cell-cycle progression. We develop PerturbMatch, a scalable statistical framework to analyze guide multiplets. Among different classes of guide multiplets, doublets and triplets recovered perturbation responses more accurately than higher-order multiplets. Across three 5000-gene Perturb-seq screens with increasing guide loading, per-cell costs decreased by up to 81% while information loss remained within 1.5-fold of the loss observed between technical replicates. In existing genome-wide Perturb-seq data, incorporating previously discarded guide multiplets increased usable cell numbers and improved statistical power. Compared with a singlet holdout set, adding guide multiplets moved signal recovery closer to the theoretical expected reproducibility. Overall, we recommend a design that intentionally includes single-guide cells, guide doublets, and guide triplets to improve cost efficiency while preserving signal recovery.
Sim, S.; Shen, L.
Show abstract
Sequence-to-function models learn regulatory features from genomic sequence, but they remain limited in their ability to predict gene-expression differences among individuals. Cell-type-specific regulatory effects may be obscured in bulk RNA sequencing, whereas paired genotype and single-cell expression cohorts remain small. We evaluated whether deconvolution of bulk RNA-seq could provide scalable cell-type-specific targets for personal-genome expression prediction. GTEx v8 bulk RNA-seq from six tissues was deconvolved with BayesPrism using single-nucleus reference profiles, producing targets across 83 tissue-cell-type contexts. Deconvolved expression agreed with matched pseudobulked GTEx single-nucleus RNA-seq, with median donor-level Pearson correlations across genes ranging from 0.53 to 0.73 by tissue. We compared genotype-feature models, regressors trained on frozen Enformer representations, and fine-tuned Enformer and Borzoi models. Across random and nonlinear-enriched gene sets, sequence-derived approaches generally outperformed genotype-feature baselines, while frozen Enformer features were competitive with end-to-end fine-tuning. For the random gene set, Fisher-averaged Pearson correlations were 0.122-0.142 for sequence-derived approaches and 0.081-0.086 for genotype-feature baselines in a coverage-aware sensitivity analysis. Model performance was positively associated with deconvolution-pseudobulk agreement for sequence-derived models (r = 0.35-0.43 across tissue-cell-type contexts), suggesting that target reliability may constrain downstream prediction. Context-specific Enformer fine-tuning did not materially out-perform a shared, combined-context strategy. These results support deconvolution as a feasible approach for generating cell-type-resolved training targets, while showing that target quality and limited cohort size remain important constraints. Frozen pretrained representations provide a computationally efficient and competitive baseline for personal sequence-to-expression modeling.
Lin, M.-J.; Shivakumar, V. S.; Langmead, B.; Human Pangenome Reference Consortium,
Show abstract
With improvements in sequencing and assembly have come many high-quality telomere-to-telomere assemblies and reference pangenomes. However, the long-read sequencing recipes needed for high quality assemblies are expensive, and out of reach for many research groups. Here we propose ImpuT2T, a method that takes an assembly produced via inexpensive HiFi sequencing reads, and uses a panel of T2T (or near-T2T) assemblies to scaffold and fill ("patch") the gaps between the HiFi contigs. Benchmarking against reference assemblies demonstrates that ImpuT2T is highly effective at patching human HiFi assemblies, consistently outperforming existing patching approaches. Moreover, we show that including more haplotypes in the pangenome improves the quality of the patched assemblies, with the greatest gains achieved using the full HPRC Release 2 pangenome.
Nikaido, I.; Shiihashi, T.
Show abstract
Background: Perturb-seq captures transcriptional responses to thousands of genetic and chemical perturbations, but does not directly resolve the cis-regulatory elements or transcription factor motifs underlying those responses. Existing approaches rely on indirect post hoc analyses or external epigenomic annotations, making it difficult to connect gene-level responses to specific regulatory element Results: We present GenPerturb, a framework that leverages pretrained sequence-to-expression models to link perturbation-induced expression changes to candidate cis-regulatory elements. By contrasting perturbation and control states, GenPerturb prioritizes regulatory regions and transcription factor motifs associated with each perturbation. The model recapitulates perturbation-dependent gene expression patterns and enables sequence-level interpretation without requiring matched chromatin data. Across multiple perturbation types, GenPerturb identifies biologically meaningful regulatory programs, including lineage-specific and signaling-associated motif activities, even when corresponding transcription factor expression changes are limited. Conclusions: GenPerturb converts gene-level expression responses from Perturb-seq into perturbation-specific, sequence-grounded cis-regulatory hypotheses. By prioritizing candidate regulatory elements and transcription factor motifs responsive to each perturbation without requiring matched chromatin data, GenPerturb enables mechanistic interpretation of transcriptional regulation and guides downstream experimental validation.
Munasinghe, M.; Read, A.; Schulz, A. J.; Brandvain, Y. J.; Springer, N. M.; Hirsch, C.
Show abstract
BackgroundStructural variants (SVs) are large insertions or deletions of DNA sequences. While less numerous than single nucleotide polymorphisms, SVs often account for a greater proportion of nucleotide differences between genomes. Their size and frequent association with repetitive sequences has historically hindered their detection, which has limited the ability to associate this variation with molecular and phenotypic trait variation. While some SVs have been linked to observable traits, it remains unclear whether such effects are rare or broadly distributed across the genome. ResultsTo test for genome-wide relationships between SVs and gene expression, we analyzed genome assemblies and transcriptomic data from 10 tissues across 26 diverse maize inbred lines. We identified SVs amongst these lines and examined variants located within the 1kb promoter region upstream of genes. Thousands of genes showed expression differences associated with promoter SVs, often in a tissue-specific manner. One common feature of these SVs was the presence of transposable element sequences. LTR retrotransposons were enriched amongst promoter SVs associated with differential expression and often reduced expression of the nearby gene. Despite widespread expression changes, we found no enrichment for specific biological functions or pathways among affected genes. ConclusionsOur findings indicate that extant TE-mediated promoter SVs play a significant role in shaping gene expression patterns across the maize genome. However, their phenotypic effects appear limited or context-dependent, suggesting that many variants may have minimal impact outside specific developmental stages or environmental conditions.
Shih, P. J.; Sanghani, Z.; Guarracino, A.; Gamaarachchi, H.; Batten, C.
Show abstract
MotivationSignal-space nanopore mappers enable real-time mapping and filtering decisions directly from raw nanopore signals. However, existing signal-space mappers are built around linear references, and using a single representative reference can introduce reference bias when the sample diverges from that reference. Pangenome reference collections can reduce this bias by representing diversity across related reference sequences, but linearreference signal mappers must treat each sequence as a separate target, redundantly storing shared sequences. Pangenome variation graphs provide a more compact representation by storing shared sequences once and encoding variants as alternative paths through the graph. Although sequence-to-graph mapping is well established for basecalled reads, existing signal-space methods do not directly use pangenome variation graphs. ResultsWe present Panomap, the first signal-space mapper that operates on pangenome variation graphs. Panomap maps raw nanopore signals to graph references, allowing signal-space mapping to use pangenome diversity while representing shared sequences once. We evaluate Panomap in three settings. First, when a single reference already maps the sample well, Panomap preserves mapping accuracy as additional reference sequences are added to the reference collection, while state-of-the-art signal-space tools regress. Second, when the exact sample strain is absent from the reference collection, Panomap benefits from adding related assemblies from the same species to the pangenome reference. Third, using a highly polymorphic locus, we show that Panomap can map reads from alleles not represented in the reference collection by using related alleles in the pangenome, with the largest gains for more divergent alleles and for decisions made from short prefixes of the read signal. In addition, Panomaps graph index scales sublinearly with pangenome collection size. Together, these results show that Panomap brings population-aware reference representation into signal-space mapping. Availability and ImplementationPanomap is open source and available at https://github.com/cornell-brg/panomap. Contactps2229@cornell.edu
Shivakumar, V. S.; Langmead, B.
Show abstract
Existing notions of pangenome coordinates rely on hard-to-compute multiple sequence alignments. On the other hand, pangenome-wide exact unique matches (multi-MUMs) can be computed efficiently, and represent conserved stretches of columns in the underlying MSA. We introduce Shredtools, which uses multi-MUMs as pangenome waypoints and allows for sophisticated queries in pangenome coordinates. Its primary query is extract, which takes an interval of one sequence and extracts the smallest window containing it that is syntenic pangenome-wide. Shredtools' extract query can extract a gene region from 476 human genomes in half a second. Other queries help to refine these results, by finding local exact matches to improve the density of multi-MUM coverage ("enhance") and by selectively discarding sequences to improve the precision of the syntenic region ("zoom"). The Shredtools web interface (available at https://vikshiv.github.io/shredtools) allows for client-side handling of extract queries with index queries handled via simple and fast HTTP Range requests, simplifying usage and enabling pangenome-scale discoveries.
Huang, Z.; Huang, R.; Han, J.
Show abstract
Sequence-to-function models predict molecular readouts directly from DNA, but recognizing a functional regulatory element is not equivalent to assigning the gene it regulates. We evaluated whether AlphaGenome deletion responses identify experimentally supported enhancer-gene relations, using a frozen K562 analysis and a primary-human-astrocyte CRISPR interference (CRISPRi) resource external to our analysis. The K562 mean contrast was positive but heavy-tailed, and exact joins showed direct collision with released Gasperini and ENCODE-rE2G resources; we therefore treated K562 as supporting evidence. In a frozen evaluation of 2,307 AstroREG relations, AlphaGenome deletion strength discriminated 133 functional relations from 2,174 well-powered nonfunctional relations (average precision 0.479, enhancer-cluster 95% confidence interval 0.394-0.561, prevalence 0.058; area under the receiver-operating-characteristic curve 0.726, 0.659-0.786). Adding AlphaGenome to distance, ABC score, enhancer length, measured expression and assay-depth context increased enhancer-grouped out-of-fold average precision from 0.396 to 0.534 and improved log loss from 0.169 to 0.150. The authors cross-fitted EGrf score was stronger alone (average precision 0.559); in a post-hoc calibration that held out both gene and enhancer folds, adding AlphaGenome increased average precision from 0.550 to 0.619 (paired enhancer-cluster increment 0.068, interval 0.023-0.115) and improved log loss from 0.143 to 0.132. This comparison had asymmetric inputs: EGrf was supervised on AstroREG labels and used local epigenomic and context features, whereas the AlphaGenome score was not fitted in this study to those labels or that feature panel but was read from a pre-existing primary-astrocyte RNA-seq output track. A post-hoc same-enhancer analysis gave conditional AUC 0.741 (0.663-0.814); a smaller same-gene analysis (34 genes, 155 relations) gave 0.701 (0.571-0.823). AstroREG labels and EGrf outputs were public before AlphaGenomes public release, so this evaluation is external to our study but not a post-release or proven-unseen benchmark. The results support complementary relation-level utility, not EGrf superiority, sequence-only deployment, causal assignment at arbitrary loci or equivalence between sequence deletion and CRISPRi.
Cai, P.; Chen, X.; Guo, J.; Wang, Y.; Yang, H.; Lin, L.; Zhang, Y.
Show abstract
Adenosine-to-inosine (A-to-I) RNA editing is a widespread post-transcriptional mechanism that diversifies the transcriptome. While ADAR enzymes catalyze this reaction, the upstream mechanisms that determine why individual adenosines are edited at markedly different efficiencies remain poorly understood. Here, we integrated multi-omics datasets from human cell lines and mouse embryonic tissues and developed machine learning models that distinguish high-and low-frequency editing sites based on local epigenetic features. Across species, tissues, and developmental stages, H3K36me3 consistently emerges as the strongest negative predictor of editing frequency, whereas the histone variant H2A.Z.1 shows a positive association. Functional validations in H2AFZ-knockdown cells reveal that H2A.Z.1 preferentially facilitates editing at high-frequency sites (76.3% of sites decreased, mean {Delta} =-0.07), whereas SETD2-knockout-mediated loss of H3K36me3 selectively derepresses editing at low-frequency sites (86.9% upregulated, [~]2.7-fold increase). These findings establish chromatin state as an upstream regulatory layer that modulates RNA editing independently of editing enzyme abundance and reveal opposing roles for H3K36me3 and H2A.Z.1 in shaping RNA editing landscapes. Together, our study provides a conceptual framework linking epigenetic regulation to post-transcriptional RNA modification, suggesting that chromatin-mediated regulation contributes to the establishment of site-specific RNA editing programs across mammalian genomes.
Sonder, E.; Aymergen, I. G.; Sun, J.; Feuvrier, A.; Schratt, G.; Gapp, K.; Bohacek, J.; Robinson, M. D.; Germain, P.-L.
Show abstract
Our understanding of the mechanisms regulating gene expression has been hampered by our limited knowledge of which transcription factors (TFs) bind where in the genome, which is highly cell type-specific. While genome-wide TF binding can be experimentally assayed for individual TFs in individual cell types, profiling the full combinations of over 1600 TFs in hundreds of cell types is beyond practical reach. In this work, we developed a streamlined platform, TFBlearner, to train TF-specific models and predict bindings based on ATAC-seq data. We focused on biologically-motivated feature engineering and harnessed TF cooperativity and binding similarity across cell types to achieve state of the art binding predictions in unseen cell types in a scalable fashion. This enabled us to generate a compendium of binding predictions for 1108 Chromatin-associated proteins, of which 960 TFs, across 43 human cell types including widely-used cell lines and 36 physiological cell types representing all major human cell lineages. We show how the models additionally provide biological insights on the TFs, and show how the binding predictions can be used in downstream tasks such as TF activity inference. Our study additionally led to the observation of high promiscuity in TF occupancy. To investigate aspecific occupancy, we characterized crowded or high-occupancy (HOT) regions across cell types, providing evidence of their functionality, and reporting important cell type-specificity. Finally, we show that, across cell types, crowded regions engage in more 3D contacts, and that most TF occupancy at crowded promoters can be explained as tethered bindings from distal regulatory elements.
Park, J.; Kang, K.
Show abstract
Differential splicing workflows usually report a {Delta}PSI point estimate and a statistical score, but these outputs do not directly state whether the RNA-seq estimate is close enough to an orthogonal validation measurement. We developed BRAID as a post-processing calibration step for splicing analyses. BRAID estimates RNA-seq {Delta}PSI from rMATS inclusion and skipping counts, retains the upstream caller evidence, and adds a 95% interval whose width is calibrated from empirical RNA-seq-to-RT-PCR residuals using split conformal prediction. The packaged differential-splicing calibrator uses a residual half-width of q = 0.341, estimated from 162 RT-PCR-validated skipped-exon events. We evaluated BRAID on three RT-PCR validation datasets covering TRA2 knockdown, mouse cerebellum versus liver, and a prostate epithelial-to-mesenchymal comparison. On the pooled common set of 139 cassette-exon events, BRAID reached 0.971 RT-PCR coverage, whereas MAJIQ, betAS, and rMATS-derived intervals reached 0.518, 0.734, and 0.633, respectively. BRAID also had the lowest pooled interval score, 0.720, compared with 2.040 for MAJIQ, 1.414 for betAS, and 1.625 for rMATS. Applying the same residual calibration to other caller outputs brought MAJIQ, betAS, rMATS, and SUPPA2 {Delta}PSI estimates close to nominal RT-PCR coverage, indicating that the gain came from interval calibration rather than from a caller-specific point estimate. In a TRA2 positive-negative validation panel, using q as a hard rMATS effect-size cutoff reduced recall, whereas using q as an interval half-width improved RT-PCR coverage. Applied to a public DM1 skeletal-muscle rMATS table, BRAID reduced 967 large-effect significant events to 68 high-confidence interval-supported events and retained known DM1 and muscle-splicing signals. BRAID provides a practical calibrated reliability layer for RNA-seq splicing studies where downstream follow-up depends on the precision of reported {Delta}PSI estimates.
Dorney, R.; Wu, S.; Hung, J. Y.-H.; Hebbard, L.; Schmitz, U.
Show abstract
Fusion transcripts contribute to cancer, inherited diseases, developmental disorders, and evolution. Long-read RNA sequencing enables direct sequencing of full-length transcripts, creating new opportunities to detect complex fusion architectures, including previously inaccessible multi-segmented fusion transcripts. However, accurate transcriptome-wide fusion detection remains challenging because existing methods struggle to distinguish genuine fusion events from technical artefacts. Here, we present a comprehensive benchmark of transcriptome-wide fusion detection using simulated datasets and transcriptomes from three cancer cell lines across Oxford Nanopore Technologies (ONT) cDNA, PCR-cDNA, and direct RNA sequencing, Pacific Biosciences (PacBio) Kinnex sequencing, Illumina short-read RNA sequencing, six long-read fusion callers, and multiple analysis strategies. False-positive fusion calls remained the dominant limitation across sequencing platforms and algorithms. Increasing sequencing depth improved recall but also amplified spurious fusion calls, whereas higher read-support thresholds improved precision at the expense of sensitivity. ONT PCR-cDNA sequencing combined with CTAT-LR-Fusion achieved the best overall balance between precision and recall, whereas JAFFAL was the only caller to reliably identify simulated tri-gene fusions. Consensus calling reduced false positives but markedly reduced sensitivity, with only one of 400 simulated fusions detected by all six callers. Breakpoint localisation emerged as a major limitation across all methods. Long-read sequencing consistently recovered more validated fusion transcripts than short-read sequencing, enabled detection of complex tri-gene fusions, and produced more biologically plausible fusion landscapes with fewer promiscuous gene partners. Collectively, our results establish the first comprehensive benchmarking framework for transcriptome-wide fusion detection, using long-read RNA sequencing, and provide practical guidance for selecting sequencing workflows and computational strategies, while identifying key priorities for future algorithm development.
Yuen, Z. W. S.; Leeder, N.; Udumanne, T.; Garvie, A.; Wong, L.; Weiss, E.; van Loon, L.; Ganley, A.; Hannan, R.; Eyras, E.; Hein, N.
Show abstract
Ribosomal RNA (rRNA) provides the structural and catalytic core of ribosomes and is encoded by ribosomal RNA genes (rDNA) arranged in tandem repeat arrays. rDNA copy number (CN) is highly dynamic, representing a clinically relevant form of structural variation, but its accurate quantification has been challenging due to its highly repetitive and GC-rich nature. Here, we present RICO (Ribosomal DNA Integrated Copy Number and Methylation Analysis), a novel computational pipeline for integrated estimation of rDNA CN and methylation using nanopore long-read sequencing. RICO leverages long sequencing reads that span entire rDNA repeats, mapped to an rDNA-augmented reference genome, and normalizes coverage using an array of single-copy genes. We show that RICO provides accurate rDNA CN estimates in simulated datasets and reproducible measurements across human samples, with strong agreement to short-read sequencing and PCR-based methods. As biological validation, RICO detects a ~40% reduction in rDNA CN in Atrx-knockout mouse cells, consistent with established effects of ATRX loss on rDNA CN, and captures detected increased total and active rDNA CN in malignant cells from a MYC-driven B-cell lymphoma mouse model, in line with prior psoralen-based chromatin studies. Applying RICO to independent human cohorts, we uncover that individuals with higher total rDNA CN consistently exhibited higher fractions of high-methylated rDNA copies, suggesting a dosage compensation mechanism that potentially maintains a similar number of active rDNA copies across individuals. Together, RICO enables integrated analysis of rDNA CN and methylation state, providing a scalable framework for investigating rDNA regulation across population and disease studies.
Bellido Molias, F.; Kudla, G.
Show abstract
Computational predictors of RNA splicing are increasingly used to interpret genetic variants and to design synthetic genes, yet they are almost always benchmarked on endogenous human sequences closely related to their training data. Whether their performance reflects genuine recognition of splicing signals, or instead exploits statistical features of natural genomes such as conservation and exon-intron composition, remains unclear. Here we benchmark eleven splicing predictors on thousands of synthetic GFP variants that are heavily recoded and dissimilar from any training data, using long-read sequencing to measure splicing directly at each position. Despite this distribution shift, modern deep-learning predictors retained strong performance, and the resulting ranking was largely stable across position-level and construct-level benchmarks. SpliceTransformer ranked highest, followed by AlphaGenome and SpliceAI. Tools that ignore long-range sequence context performed substantially worse, largely because they assign high scores to many non-spliced positions. This ranking broadly agrees with benchmarks on endogenous variants, indicating that the leading models capture transferable, sequence-intrinsic determinants of splicing. We further provide a unified calibration that maps each predictor's scores onto the measured fraction of spliced reads, allowing scores to be interpreted as splicing outcomes and compared directly between tools. Our results show that current deep-learning models generalise beyond natural genomes and provide a practical framework for splicing-aware sequence design.
Asiaee, A.; Bombina, P.; McGee, R. L.; Reed, J.; Abrams, Z. B.; Abruzzo, L. V.; Coombes, K. R.
Show abstract
BackgroundCorrelation is the dominant input to co-expression module discovery and miRNA-target inference. Both rely on an implicit assumption: a Pearson coefficient pooled across heterogeneous samples, whether tissues, cancer types, or cell types, estimates one biologically meaningful quantity. Simpsons paradox makes this assumption fragile in principle, since between- group mean shifts can dominate or reverse within-group associations. How often this happens in real transcriptomic data has not been quantified. ResultsAcross 8,890 TCGA tumors from 31 cancer cohorts and 23,170,038 miRNA-mRNA pairs, 94.8% of pairs showed both positive and negative within-cohort correlations. Restricting to the high-variance domain of one million pairs, 13.3% of pooled correlations with |rglobal|[≥]0.2 reversed against the within-cohort majority at sign tolerance{varepsilon} = 0.05. Heterogeneity was the rule rather than the exception (median I2 = 0.86, IQR 0.80-0.90), and 99.5% of pairs rejected equal correlation across cohorts at FDR < 0.05. Of 692,770 experimentally validated miRTarBase v10 targets measurable in our data, only 0.9% were uniformly negative across cohorts. The pattern recurred across modalities. In GTEx, 21.0% of pooled signs disagreed with the tissue majority, and 23.5% of pairs flipped sign after tissue-mean removal. In 10x PBMC scRNA-seq, 13.1% of gene-gene correlations flipped after cell-type-mean removal; in CITE-seq, 37.9% of protein-RNA pairs flipped under a joint WNN partition of cells. Refining context reduced reversal, though by how much depended on the partition: within BRCA, 5.5% of pairs reversed under molecular PAM50 subtypes versus 0.35% under clinical IHC receptor status, and refining T cells into transcriptome-defined subtypes cut PBMC reversal from 11.8% to 0.13%. ConclusionsA single pooled correlation coefficient can invert direction relative to its within-context constituents at rates that are not negligible. Correlations should be reported with their context: the within-context distribution, a heterogeneity statistic, and a diagnostic that separates between-context mean shifts from within-context association. We provide a small R interface that computes these summaries.